Free your Camera: 3D Indoor Scene Understanding from Arbitrary Camera Motion
نویسندگان
چکیده
Many works have been presented for indoor scene understanding, yet few of them combine structural reasoning with full motion estimation in a real-time oriented approach. In this work we address the problem of estimating the 3D structural layout of complex and cluttered indoor scenes from monocular video sequences, where the observer can freely move in the surrounding space. We propose an effective probabilistic formulation that allows to generate, evaluate and optimize layout hypotheses by integrating new image evidence as the observer moves. Compared to state-of-the-art work, our approach makes significantly less limiting hypotheses about the scene and the observer (e.g., Manhattan world assumption, known camera motion). We introduce a new challenging dataset and present an extensive experimental evaluation, which demonstrates that our formulation reaches near-real-time computation time and outperforms state-of-the-art methods while operating in significantly less constrained conditions. Figure 1 shows a pictorial representation and a schematic diagram of the whole process. Sparse 3D reconstruction. As the observer moves in the surrounding environment, we first pre-process sequences with a localization and sparse 3D reconstruction algorithm. In our experiments we compare two such approaches: a real-time implementation of the Monocular V-SLAM approach proposed in [4] and the non-real-time VisualSfM [6]. These 3D reconstructions are in general noisy and sparse. Candidate layout components. The second step consists of generating a higher level representation of the 3D points estimated in the preprocessing phase. Several types of geometrical primitives are suitable for this purpose. In our case, we believe a piecewise planar representation is the most appropriate for indoor scene representation. We fit a large number of planes to the 3D points so as to generate a large number of (potentially inaccurate) candidates of layout components, i.e. walls, floor, ceiling. In our experiments we implemented an Iterative RanSaC plane fitting procedure, which we optimized for indoor scenes by allowing peripheral fitted points to be re-injected in the iteration process, since these points potentially lay on the intersection of two planes. Layout Parametrization While much prior work has leveraged the Manhattan world assumption, we believe this is a limiting hypothesis. To overcome this limitation, in this paper we adopt a representation similar to [5] (sometimes referred to as Soft Manhattan), which makes the following assumptions about the environment: i) ground plane and ceiling are parallel; ii) walls are only constrained to be orthogonal to the ground plane (and ceiling); iii) there can be any number of walls and each wall can be displaced at any angle with respect to other walls. Layout estimation. In the last step, constituting the core of our proposed inference engine, we generate layout hypotheses as random combinations of candidate layout components. Each layout hypothesis is evaluated at each time frame by measuring its compatibility with observations (e.g. image points and lines) and geometrical constraints across frames. During this process, each layout is “perturbed” by locally adjusting, optimizing, merging or splitting layout components. There are different approaches to manage sets of hypotheses. In this paper we choose to integrate our probabilistic framework within a particle filter structure. This choice allows to explicitly formulate the problem in a parallel-computing oriented fashion (particles are independent from each other), which can lead to high efficiency gains in computation time. The output of the optimization procedure is an estimation of the 3D scene layout, which is obtained by selecting the layout hypothesis with the best set of layout components. Figure 1: 3D scene layout estimation process. The video sequence is first processed to obtain camera localization and sparse 3D point cloud reconstruction. Layout components (e.g. floor, ceiling, walls) are generated from the sparse 3D points and combined to generate layout hypotheses. Each layout hypothesis is evaluated and optimized by incorporating new image evidence. The final 3D scene layout is represented by the hypothesis that better describes the scene.
منابع مشابه
A Statistical Geometric Framework for Reconstruction of Scene Models
This paper addresses the problem of reconstructing surface models of indoor scenes from sparse 3D scene structure captured from N camera views. Sparse 3D measurements of real scenes are readily estimated from image sequences using structure-from-motion techniques. Currently there is no general method for reconstruction of 3D models of arbitrary scenes from sparse data. We previously introduced ...
متن کاملGenerating Free Viewpoint Images from Mutual Projection of Cameras
In augmented reality, accurate geometric adjustment of real scene and virtual 3D models is important. In this paper, we propose a new method for generating arbitrary views of 3D motion events accurately by using the mutual projections between user’s cameras and cameras around the user. In particular, we show that the trifocal tensors computed from the mutual camera projections can be used effic...
متن کاملModel–Based 3D Scene Analysis from Stereoscopic Image Sequences
An approach for the modelling of complex 3D scenes like outdoor street views from a sequence of stereoscopic image pairs is presented. Starting with conventional stereoscopic correspondence analysis a 3D model scene with true 3D geometry is generated. Not only the scene geometry but also surface texture is stored within the model. 3D camera motion can be estimated directly from the image sequen...
متن کاملExploring High-Level Plane Primitives for Indoor 3D Reconstruction with a Hand-held RGB-D Camera
Given a hand-held RGB-D camera (e.g. Kinect), methods such as Structure from Motion (SfM) and Iterative Closest Point (ICP), perform poorly when reconstructing indoor scenes with few image features or little geometric structure information. In this paper, we propose to extract high level primitives– planes–from an RGB-D camera, in addition to low level image features (e.g. SIFT), to better cons...
متن کاملPedestrians Tracking in a Camera Network
With the increase of the number of cameras installed across a video surveillance network, the ability of security staffs to attentively scan all the video feeds actually decreases. Therefore, the need for an intelligent system that operates as a tracking system is vital for security personnel to do their jobs well. Tracking people as they move through a camera network with non-overlapping field...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2013